Papers with open-source tool

13 papers
LM-Debugger: An Interactive Tool for Inspection and Intervention in Transformer-Based Language Models (2022.emnlp-demos)

Copied to clipboard

Challenge: Transformer-based language models (LMs) are opaque and unexplained, causing problems for endusers and developers who wish to debug or fix their behaviour.
Approach: They propose an interactive debugger tool for transformer-based LMs that provides a fine-grained interpretation of the model's internal prediction process and a powerful framework for intervening in LM behavior.
Outcome: The proposed tool provides a fine-grained interpretation of the model's internal prediction construction process, and a powerful framework for intervening in LM behavior.
Loki: An Open-Source Tool for Fact Verification (2025.coling-demos)

Copied to clipboard

Challenge: Loki is an open-source fact-checking tool designed to address the growing problem of misinformation.
Approach: They propose a tool that breaks down the fact-checking task into five steps . they propose LOKI, which offers a semiautomated, human-in-the-loop approach .
Outcome: a new open-source tool is designed to address the growing problem of misinformation . the tool breaks down the fact-checking task into five steps to assist human judgment .
NameTag 3: A Tool and a Service for Multilingual/Multitagset NER (2025.acl-demo)

Copied to clipboard

Challenge: NameTag 3 is an open-source tool and cloud-based web service for named entity recognition.
Approach: NameTag 3 is an open-source tool and cloud-based web service for named entity recognition.
Outcome: NameTag 3 achieves state-of-the-art on 21 test datasets in 15 languages . available as command-line tool and as cloud-based service, enabling use without local installation .
YEDDA: A Lightweight Collaborative Text Span Annotation Tool (P18-4)

Copied to clipboard

Challenge: Existing annotation tools do not consider post-annotation quality analysis due to inter-annotator disagreement.
Approach: They propose a lightweight but efficient open-source tool for text span annotation that can be used for collaborative user annotation and administrator evaluation and analysis.
Outcome: The proposed system reduces the annotation time by half compared with existing tools and the time can be compressed by 16.47% through intelligent recommendation.
A Multiscale Visualization of Attention in the Transformer Model (P19-3)

Copied to clipboard

Challenge: Various tools have been developed to visualize attention in NLP models, ranging from attention-matrix heatmaps to bipartite graph representations.
Approach: They propose an open-source tool that visualizes attention at multiple scales and provides a unique perspective on the attention mechanism.
Outcome: The proposed model outperforms OpenAI GPT-2 and BERT on several language modeling benchmarks.
SLTEV: Comprehensive Evaluation of Spoken Language Translation (2021.eacl-demos)

Copied to clipboard

Challenge: Spoken Language Translation (SLT) evaluation of machine translation (MT) quality has been investigated for decades.
Approach: They propose an open-source tool for assessing machine translation (MT) quality based on time-stamped transcripts and reference translations.
Outcome: The proposed evaluation tool is based on time-stamped transcripts and reference translations into a target language.
kogito: A Commonsense Knowledge Inference Toolkit (2023.eacl-demo)

Copied to clipboard

Challenge: kogito provides an intuitive and extensible interface to interact with natural language generation models.
Approach: They propose to use kogito to generate commonsense inferences from text . they use a standardized API for training and evaluating knowledge models .
Outcome: The proposed tool provides an intuitive and extensible interface to interact with natural language generation models.
AlignFix: A Tool for Parallel Corpora Augmentation and Refinement (2026.eacl-demo)

Copied to clipboard

Challenge: High-quality datasets are crucial for training effective state of the art machine translation systems, but they can be noisy and degrade performance.
Approach: They propose an open-source tool for augmenting data, identifying and correcting errors in parallel corpora.
Outcome: The tool extracts consistent phrase pairs, enabling targeted replacements that can improve the dataset quality.
SummVis: Interactive Visual Analysis of Models, Data, and Evaluation for Text Summarization (2021.acl-demo)

Copied to clipboard

Challenge: despite advances in abstractive text summarization, the true performance and failure modes of modern neural models are not yet fully understood due to the black-box nature of neural models and unmanageable scale of recent datasets for manual analysis.
Approach: They propose an open-source tool for visualizing abstractive summaries that enables fine-grained analysis of models, data, and evaluation metrics associated with text summarization.
Outcome: The proposed tool can identify the shortcomings and failure modes of state-of-the-art summarization models.
SynKB: Semantic Search for Synthetic Procedures (2022.emnlp-demos)

Copied to clipboard

Challenge: SynKB is an open-source, automatically extracted knowledge base of chemical synthesis protocols.
Approach: They propose to make SynKB available as an open-source tool for chemists . synKB supports more flexible queries about reaction conditions .
Outcome: The proposed open-source tool has higher recall and high precision than proprietary chemistry databases.
The AI Committee: A Multi-Agent Framework for Automated Validation and Remediation of Web-Sourced Data (2026.eacl-demo)

Copied to clipboard

Challenge: largelanguage models (LLMs)-powered web agents can be useful for research in areas such as social science, public health, and economics.
Approach: They propose a model-agnostic multi-agent system that auto-mates the process of validating and remediatingweb-sourced datasets.
Outcome: The proposed system outperforms baseline approaches and achieves datacompleteness and precision up to 73.3%.
LLMDet: A Third Party Large Language Models Generated Text Detection Tool (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing detection tools rely on access to LLMs and can only distinguish between machine-generated and human-authored text.
Approach: They propose a model-specific, secure, efficient, and extendable detection tool that can source text from specific LLMs.
Outcome: The proposed tool can source text from specific LLMs, such as GPT-2, OPT, LLaMA, and others.
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models (2025.coling-main)

Copied to clipboard

Challenge: Large language models excel in various language tasks, while large multimodal models effectively handle visual-language problems.
Approach: They propose to use a multimodal multimodal model evaluation benchmark to evaluate model performance in Chinese K12 classrooms.
Outcome: The proposed model evaluation tool is integrated with the CMMaTH dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations